Papers with speech resynthesis
textless-lib: a Library for Textless Spoken Language Processing (2022.naacl-demo)
Copied to clipboard
Eugene Kharitonov, Jade Copet, Kushal Lakhotia, Tu Anh Nguyen, Paden Tomasello, Ann Lee, Ali Elkahky, Wei-Ning Hsu, Abdelrahman Mohamed, Emmanuel Dupoux, Yossi Adi
| Challenge: | Textless spoken language processing is an exciting area of research that promises to extend applicability of the standard NLP toolset onto spoken language and languages with few or no textual resources. |
| Approach: | They introduce textless-lib, a PyTorch-based library that provides textless spoken language processing tools. |
| Outcome: | The proposed library significantly simplifies research in the textless setting and will be a handful for speech researchers and the NLP community at large. |
EmphAssess : a Prosodic Benchmark on Assessing Emphasis Transfer in Speech-to-Speech Models (2024.emnlp-main)
Copied to clipboard
| Challenge: | EmphAssess evaluates speech-to-speech models' ability to encode and reproduce prosodic emphasis across a change of speaker and language. |
| Approach: | They propose a prosodic benchmark to evaluate the ability of speech-to-speech models to encode and reproduce prosodic emphasis. |
| Outcome: | The proposed model can encode and reproduce prosodic emphasis across speech inputs and outputs . EmphaClass classifies emphasis at the frame or word level . |